Papers by John P. McCrae
Linghub2: Language Resource Discovery Tool for Language Technologies (2022.lrec-1)
Copied to clipboard
| Challenge: | Linghub is a platform for language resources that can be used to find and retrieve data . the platform is based on a popular open source data management system, DSpace . |
| Approach: | This work describes a rejuvenation and modernisation of the 2015 platform into using a popular open source data management system, DSpace, as foundation. |
| Outcome: | Linghub2 1 aims to help language resources and technology users find and retrieve relevant data . the new platform, Ling hub2, contains updated and extended resources and more languages offered . |
Teanga: A Linked Data based platform for Natural Language Processing (L18-1)
Copied to clipboard
| Challenge: | Using linked data, we can use many NLP services from a single interface . integrating components within a development model is endemic to software development . |
| Approach: | They propose a linked data based platform for natural language processing that uses linked data to define the types of services input and output. |
| Outcome: | The proposed platform is easy to install and run, easy to use and able to run multiple NLP tasks from one interface. |
Cross-lingual Sentence Embedding using Multi-Task Learning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing multilingual sentence embedding models require large parallel corpora to learn efficiently, limiting their scope. |
| Approach: | They propose a sentence embedding framework based on an unsupervised loss function . they capture semantic similarity and relatedness between sentences using a multi-task loss function. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on STS, BUCC and Tatoeba benchmarks and on a monolingual benchmark. |
Contextual Modulation for Relation-Level Metaphor Identification (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to identifying metaphors in text ignore context where metaphor occurs . existing approaches focus on word-level identification without explicitly modelling interaction between metaphor components . |
| Approach: | They propose a method for identifying relation-level metaphoric expressions of certain grammatical relations based on contextual modulation. |
| Outcome: | The proposed architecture achieves state-of-the-art results on benchmark datasets. |
Unsupervised Deep Language and Dialect Identification for Short Texts (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for identifying closely related short texts are unsupervised . however, performance is poor for unsupervised methods for short texts . |
| Approach: | They propose a method which can learn sentence embeddings and cluster assignments from short texts. |
| Outcome: | The proposed method outperforms state-of-the-art methods in supervised settings . it can learn sentence embeddings and cluster assignments from short texts . |
A supervised approach to taxonomy extraction using word embeddings (L18-1)
Copied to clipboard
| Challenge: | a recent evaluation of a method for organizing texts into a hierarchy showed that it did not outperform a baseline. |
| Approach: | They propose a method that uses supervised learning to combine multiple features with a support vector machine classifier including the baseline features. |
| Outcome: | The proposed method outperforms the baseline method and provides stronger method for identifying taxonomic relations than previous methods. |
Suggest me a movie for tonight: Leveraging Knowledge Graphs for Conversational Recommendation (2020.coling-main)
Copied to clipboard
| Challenge: | Recent studies show that knowledge graphs are incomplete since they do not contain all factual information present on the web. |
| Approach: | They propose to use knowledge graphs to improve the performance of conversational recommender systems by incorporating pre-trained embeddings from subgraphs and positional embeddments into their models. |
| Outcome: | The proposed method improves by 5.62% over the state-of-the-art method on multiple metrics on the recommendation task. |
Towards the Construction of a WordNet for Old English (2022.lrec-1)
Copied to clipboard
Fahad Khan, Francisco J. Minaya Gómez, Rafael Cruz González, Harry Diakoff, Javier E. Diaz Vera, John P. McCrae, Ciara O’Loughlin, William Michael Short, Sander Stolk
| Challenge: | In this paper we discuss our preliminary work towards the construction of a WordNet for Old English, taking our inspiration from other similar WN construction projects for ancient languages such as Ancient Greek, Latin and Sanskrit. |
| Approach: | They propose to use a legacy Old English dictionary to build a WordNet for Old English using a lexicographic resource and the naisc system to automatically compile a provisional version of the WordNet. |
| Outcome: | The proposed OldEWN will be based on lemmas and definitions extracted from a legacy Old English dictionary and will be automatically compile and enriched by experts using the naisc system. |
MHE: Code-Mixed Corpora for Similar Language Identification (2022.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of code-mixed data-sets is presented for similar language identification . the data-settings are based on a more-resourced minority language, Magahi . |
| Approach: | They propose a Magahi-Hindi-English code-mixed corpus for similar language identification . they discuss the complexity of the data-set and provide a few baselines . |
| Outcome: | The proposed corpus provides a language id at two levels: word and sentence. |
MaCmS: Magahi Code-mixed Dataset for Sentiment Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | Sociolinguists and psychologists have been studying these variations in the lexicons and the language from the 50's . code-mixing is a popular method for understanding people's emotions and attitudes towards various subjects, but low-resourced languages often have a mix of scripts and languages. |
| Approach: | They introduce a new sentiment data, MaCMS, for Magahi-Hindi-English code-mixed language, where Magai is a less-resourced minority language. |
| Outcome: | The proposed dataset is the first Magahi-Hindi-English code-mixed dataset for sentiment analysis tasks. |
Automatic Enrichment of Terminological Resources: the IATE RDF Example (L18-1)
Copied to clipboard
| Challenge: | a recent paper aims to automate the maintenance of terminological resources. |
| Approach: | They propose automatic approaches to maintain and increase lexical coverage of knowledge bases by using machine translation and multilingual word sense disambiguation. |
| Outcome: | The proposed approach outperforms the existing methods with random sentences in most languages . |
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)
Copied to clipboard
| Challenge: | a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes . |
| Approach: | They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones . |
| Outcome: | The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions. |